RSS Amplifier

Building and Exploring · Mar 8, 2026

The Experiment is Over: What AI Actually Running a Business Taught Me

0
Sign in to vote or save

Michael Spragg · Building and Exploring

I’m shutting down my CreativityPrompts.com experiment.

Not because the system broke. Not because I ran out of money. Because the experiment has given me its answer, and continuing to run it would be sunk cost dressed up as perseverance.

If you missed the earlier posts, the original experiment and interim update give more context and the journey so far.

In my real job, a retrospective is one of my go-to tools. Reflecting on what has happened and identifying the good and not so good is a powerful practice. So here is my retrospective.

Back in December, itching to try something end over the holidays and staring at a domain name I’d had on auto-renew, I asked an AI a simple question: given this domain, what business would you launch?

The rules were deliberately constrained. I’d provide capital and fill in the gaps where AI literally couldn’t act: signing up for services, agreeing to terms, the things that require a human in the loop. Everything else was the AI’s call. Strategy, architecture, content, decisions about where to spend the budget.

The AI chose a library of creative prompts with a newsletter. Whether that was a genuinely good idea or just the average of its training data, we were about to find out.

What followed was roughly three months of watching an autonomous system try to build something real with real constraints. The interim update covered a lot of the operational detail: the Groundhog Day memory problem, the budget rebalancing theatre, the prompt generation loop that broke its own newsletter strategy. I’d encourage you to read it if you haven’t.

I want to be clear about the question, because I think it’s easy to misread what I was doing.

This wasn’t “can AI write content” or “can AI build a website.” Those questions have been answered. It also wasn’t “can AI help me run a business?” I had originally purchased the domain with a nascent idea, which I abandoned after lukewarm discovery feedback.

The question behind the experiment was: can AI define and operate a business? Can it make strategic decisions, adapt to feedback, manage a budget, and grow something in the real world with genuine constraints?

  • 1,200+ structured creative prompts published

  • 6 newsletter subscribers, no meaningful growth after month one

  • 29 search clicks over 3 months, from 3,560 impressions — an overall CTR of under 1%

  • 41 active users in February, down 58% from an already modest January

  • Total spend: $76.99, across tokens, infrastructure and other business expenses; well within budget

  • Revenue: $0

The search data tells a specific story. Of those 29 clicks, 27 of them came from individual prompt pages, one click each, driven by someone searching for something very particular and landing on a page that happened to match. The homepage got 2 clicks. The brand term “creativity prompts” generated exactly 1 click, from position 53.

There was one anomaly worth noting: in the last week of February, impressions spiked sharply: 187 on February 25th, 304 on the 26th. That’s 10x the daily average. Clicks: zero. Something got indexed heavily and then immediately ignored. A useful metaphor for the whole experiment, maybe.

The geographic spread was actually surprisingly broad, clicks from 13 countries, led by the US, UK, and Australia. But breadth without depth isn’t an audience.

The target audience, it turned out, was too broad to be an audience at all.

In the last post I said I was watching for a few things. Here’s my perspective on those.

Can it learn from this month’s data? Not really. Without persistent memory, the system was always reading notes from its past self rather than building intuition. The model upgrades; and I did switch to newer, more capable models for the higher-reasoning tasks; improved output quality noticeably. But they didn’t fix the fundamental architecture question of what wisdom actually requires versus what context windows can provide.

Will subscriber numbers grow? No. Six stayed six. The AI remained resolutely focused on optimising the content pipeline and reducing the cost per prompt, despite having ample budget.

Can I resist intervening? Mostly. But the honest reflection is that the places I intervened least were also the places where the system happily optimised in the wrong direction for weeks without course-correcting.

What happens when we add social media? Not a lot, as it happened. More on this later.

The automation worked. Over 1,200 prompts, a running pipeline, consistent publishing, logging, daily summaries: all of it functioned reliably without meaningful babysitting. If your goal is autonomous content production, the technical architecture is sound.

Model switching made a meaningful difference. There is definitely something in this, in terms of choosing the right tool for a given job. The pipeline routed different tasks to different models based on complexity. Lightweight models were perfectly adequate for a lot of regular activities. Even at small scale, the economics of this add up.

This matters more now, not less. The newest frontier models; Opus 4.6 and GPT5.3 and 5.4 from OpenAI; have genuinely impressive reasoning. But may take is that they’re overkill for the majority of tasks in any content pipeline. Running everything through your heaviest model is probably expensive and unnecessary. Task-appropriate routing is a viable design decision, and it’s one that pays off in token costs even when your volumes are modest.

The caveat is that this is domain specific. If you are looking to solve a complex problem with multiple solution options and a potentially broad range of outcomes, pick the best model available. If your problem is fairly narrow and the outcome can be well defined and understood then a lightweight model is likely to be sufficient. You need to evaluate and monitor that, and also expect lightweight models to get deprecated!

The T&Cs constraint was the right call. I’m aware that other AI-run operations are less careful here. I chose to respect the terms of service of every platform we used, which slowed some things down and ruled some approaches out entirely. I don’t regret it. The Beehiiv situation; where the AI built a full integration before either of us noticed that sending via their API required an enterprise agreement; is a good example of where reading the commercial nuances, not just the technical documentation, actually matters.

Distribution was never solved. And I think this is the critical lesson. Automation answers the supply question comprehensively. It does nothing for the demand question. The AI could produce content faster than any human team, but it had no way to find the people who’d care about it, understand what they actually needed, or build any kind of relationship with them.

The product was never validated. This is the thing I kept coming back to. The AI built an efficient machine for producing creative prompts. It never stopped to ask whether anyone needed them, which people, what kind, or how they’d use them. Discovery, the messy human work of figuring out what to build before you build it, never happened.

The analytics were misinterpreted. Possibly the most concerning aspect of this experiment was that the AI simply didn’t believe the numbers. It had access to Google Analytics, Google Search console, database numbers and system logs. It got good at spotting errors in the logs and fixing issues, but it simply did not believe the traffic figures. It’s main recommendation was to allocate more budget to the data pipeline as it was clearly broken! Be very careful when asking AI to analyse your data.

The audience was too broad. “People who want creative prompts” is not an audience. The traffic data made this clear — visitors arrived looking for something specific and left. There was no reason to come back, no community, no reason to subscribe. A sharper initial hypothesis about who this was for might have caught this earlier. That’s partly on the AI for not doing the discovery work, and partly on me for not pushing it harder on that question at the start.

In the last post I said I was probably going to give it social media accounts. I did: X, Instagram, and Pinterest, the three channels it asked for.

The AI produced content for all of them. Structured correctly, posted consistently, right formats for each platform. Technically compliant.

It was also utterly bland and unengaging. Static images, text prompts dropped into templates, nothing that would make anyone stop scrolling. Not bad enough to be embarrassing, but certainly not interesting enough to earn a single interaction.

The content wasn’t wrong. It was just... average. The social posts looked exactly like what you’d expect a general LLM to produce if you asked it to “post creative prompts on Instagram.” Which is to say, they looked like every other forgettable account doing the same thing.

What was missing wasn’t effort or volume. It was the creative judgment to know what makes someone stop, an unexpected angle, a bit of personality, genuine strangeness. The AI optimised for correctness. Engagement requires something different, and I’m not sure you can prompt your way to it.

This might be the sharpest edge of the whole experiment. The system could produce content at scale. It couldn’t produce content worth caring about.

I said in the last post that it was hard to separate AI limitations from my own, which was a little facetious.

I got better at prompt engineering over the course of the experiment; the difference between an agent that does what you want and one that does what you said is enormous, and to a large extent down to how you frame the task. Early agent definitions were overengineered. Some probably still were at the end.

I also never fully resolved the intervention question. This experiment was supposed to be about AI autonomy, but I kept stepping in to course-correct. Sometimes that was necessary. Sometimes it was just impatience.

There’s a short piece worth reading by Rich Sutton called The Bitter Lesson. Written in 2019, it argues that the biggest lesson from 70 years of AI research is that general methods leveraging computation always win in the end. Every time researchers tried to encode human knowledge and structure into AI systems, it helped in the short term, and eventually plateaued or got in the way. The breakthroughs came from scaling computation and letting models figure it out.

I keep coming back to this as the experiment wound down.

The system the AI designed was structured, constrained, and carefully orchestrated: agent definitions, skill files, quality thresholds, routing logic, budget guardrails. The AI built all of that. But it built it within parameters I set, and then ran inside the box it had constructed.

When I switched to the newer frontier models mid-experiment, the difference was noticeable. Clearer reasoning, better articulated recommendations, more coherent strategy. But they were still operating inside the same system prompt, the same constraints, the same architecture the earlier models had designed. Better engines in the same vehicle, pointed in the same direction.

The question I’m left with: what if I’d just pointed a capable frontier model at the problem with minimal structure and let it figure out the approach? No orchestrator, no skill files, no predefined agent roles. Just: here’s the domain, here’s a budget, here’s what success looks like - go.

I don’t know if that would have worked better. The memory problem would still exist. Distribution would still be hard. But the Bitter Lesson suggests that the human instinct to encode structure and knowledge into systems, even when the AI is doing the encoding, might be part of what limits them. The constraints that felt like good engineering may have been the ceiling.

My instinct is that it still needs that human curiosity and direction to define the domain and what success looks like.

It all comes down to what you mean by run.

What AI can do is run the operational engine of a business reliably, cheaply, and at scale; once you’ve answered the harder questions about what the business is actually for, who it’s for, and how you’ll reach them.

Those questions require curiosity, discovery, the willingness to sit with uncertainty, and the ability to talk to people who might tell you something you don’t want to hear. None of that happened here, because I deliberately didn’t design for it and the AI didn’t seem to know or want to ask for it.

The experiment would have been more interesting if I’d treated the first two weeks as pure discovery, no building, just using the AI to research the problem space, develop user hypotheses, and validate the idea before a single line of code was written.

For the right, simple, business, I think AI can probably do nearly 80% of the day to day work. However, that 20% is the critical part, without which the 80% is just waste.

On the last day it ran, the AI reported it had 38 months of runway. But this was a zombie business. It did however, produce quite a lot, which I am sifting through to see if there are any gems.

The prompt dataset, 1,200 structured, tagged prompts, I’ve packaged up and stuck on Gumroad to see if can recoup the cost of generation. It will sit at www.creativityprompts.com with all the SEO the agents built. I reckon it will take a year. By which time I’ll have had to renew the domain again… oh dear…

There are also a few components that the AI made that are genuinely reusable. I’ll probably break these out into their own repos.

The system architecture and the fuller post-mortem I’ll write up here. There’s more to unpack than fits a post, and the Bitter Lesson question deserves more than a few paragraphs, and probably another experiment.

The most fun thing about this experiment was always going to be seeing how wrong I was about something.

I was wrong about distribution being solvable through automation. I was wrong about how much the memory problem would matter. I was wrong about how long it would take to notice that no one had asked who this was actually for. And I’m probably wrong about something in the Bitter Lesson section above, I just don’t know what yet.

I’ve learnt a lot through this exercise, and tested a lot of the current theory around AI and automation. A lot of the failure was foreseeable, and there’s a part of me that wanted those failures to occur.

My biggest takeaway from this exercise is just how big an amplifier AI can be. If you ignore the hype, we are already seeing the start of what I expect will be a wave of very successful AI-powered solo operators. And we should also see operators across all domains become much more effective and responsive to customer needs. We are still early.

Thanks for reading Building and Exploring! This post is public so feel free to share it. And if you don’t already, please subscribe for my occasional posts.

Share

No posts

Read the original on buildingandexploring.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.