RSS Amplifier

Everscale · Jun 9, 2026

Deepfake Accounting, Anyone?

0
Sign in to vote or save

Gary Turner · Everscale

Note: Despite some expressed scepticism in this article about how LLMs are being deployed, I am “long” on AI and LLMs and convinced that they will be transformational and disruptive across many industries and settings.

I’m a fan of analyst Benedict Evans’ pragmatic takes on most things, and last week he shared his view of where we are, which aligns with my outlook:

“We’re in 1997 for AI— it’s as big a deal as the internet or mobile, and only as big a deal as the internet or mobile. We’re at the stage where most stuff kind of doesn’t work yet, most of what people will build hasn’t been built, and it’s not clear how any of it will work when it does.” — Benedict Evans, May 2026.

Every startup begins with a central thesis. To attract capital, that thesis usually needs to be validated by data and market insights showing an identifiable, large, and obtainable customer base that feels enough pain to pay for a solution that will result in sustainable profitability for the startup.

In January, I wrote that the era of AI-assisted coding was about to change the economics of software development and, with it, this start-up equation:

“We will be awash in niche, good enough software, many of which will go nowhere or fail. We are witnessing the democratisation of engineering, where the barrier to entry drops from millions of dollars to a good idea.” — Pattern Recognition, January 2026

Six months later, the investor pitches are flooding in. I now see prototypes every week that would have been impossible to build just half a year ago. On the surface, this is a victory that demonstrates the democratisation of engineering is real.

But there is a concerning parallel: while the barrier to building has dropped, the barrier to strategy remains as high as ever. I’m seeing founders with working prototypes who have done zero market analysis and have no go-to-market plan. They have a vibe-coded product, a founder’s hunch, and very little else.

In some respects, it’s reminiscent of 2008 when the iPhone App Store first launched, and the Top 10 lists were littered with novelty apps:

  • Digital pints of beer that “emptied” as you tilted the phone.

  • Lightsaber simulators with swooshing sound effects.

  • The ubiquitous “pull-my-finger” fart apps.

AI-assisted coding is doing for development what the App Store did for distribution. It has flattened the cost of version 1.0, but it has left the cost of everything else unchanged. In fact, the non-product work —positioning, distribution, and customer acquisition—now carries more weight because the market is suddenly much noisier.

The novelty apps of 2008 died because they weren’t real businesses. They were jokes that stopped working after an iOS update or two because their hobbyist developers couldn’t justify the upkeep. Bringing a vibe-coded app to market in a fortnight is one thing; preventing it from decaying into a liability is another entirely.

And so, it was with this observation fresh in my mind that I read Ted Chiang’s article, "No, Artificial Intelligence is Not Conscious," in The Atlantic last week, in which he explored the philosophical question of whether LLMs can manifest consciousness. Spoiler alert: if the title hasn’t already let the cat out of the bag, he argues against. Chiang argues that LLMs, on account of their guess-the-next-word-in-the-sequence manner of artificial thinking, are no more capable of consciousness than a Microsoft Word document is.

Claude might convincingly simulate an engaging conversation with a human in a customer service encounter, and Gemini can role-play and emulate some semblance of consciousness, but as powerful as they are — and language alone seems to be extending the magic trick far beyond what even LLM’s inventors might have predicted at the outset — ultimately they’re bound by the immutable limitations of their design and are therefore incapable of being conscious in any true sense of the meaning.

Instead, Chiang punches through to the other side of the question by arguing that in the age of LLMs, “we need to regard text as a Deepfake medium”. Until now, deepfakes have been exclusively associated with the AI-assisted production of convincing but fake videos and imagery, and some of the earliest deepfakes even predate LLMs.

This is a profound insight, and I think it’s worth pulling on this thread and considering whether the characteristics of deepfakes can rationally be extended not only to AI-generated text but also applied to AI-assisted or vibe-coded apps, as well as to the manner in which apps increasingly embed and encapsulate probabilistic LLM-built capabilities and functionality.

Which brings us to the Uncanny Valley. If a deepfake is a very convincing facade but with no structural integrity behind it, then AI-assisted coding may well be producing a generation of “Deepfake Startups.”

It looks like a duck, swims like a duck, and quacks like a duck, but it’s not a duck?

These products look almost exactly like fully built business applications. They move data plausibly. They feel real to a non-expert (and, in some cases, even to their creators) but aren’t built on deep product thinking or domain expertise; rather, they are facsimiles of software.

And even if such products were capable of withstanding rigorous functional and technical scrutiny and evaluation, then the associated absence of broader market insight, strategy, a coherent business plan or basic GTM thinking, in my view, still tags them as deepfakes, since it’s likely they’ll never reach commercial viability and find a market, despite exhibiting the swagger and confidence that they could.

As an aside: being an early-stage investor, I’m now wondering if I need to devise a version of the Voight-Kampff Test from the 1982 movie Blade Runner: a series of interview questions in a version of a polygraph test employed by Deckard, the film’s lead, portrayed by Harrison Ford, designed to detect whether a person is human or android.

In accounting contexts, this surface-and-nothing-behind-it problem may introduce more serious implications that most of the market has not yet noticed. Some of these products do not merely look like real software — they use LLMs to perform some of the core functions of the product itself, which in the domain of accounting, and specifically making determinations about categorising transactions against a set of accounting principles and the Chart of Accounts (the strict accounting taxonomy for any business), means that important elements of the accounting logic are not rule-based and deterministic but probabilistic, generating outputs through pattern matching rather than through the application of accounting rules to source evidence, and in a way that is invisible to the user and undetectable without complete re-derivation of the figures from source data and paperwork.

I’m on the board at Mazuma, a fast-growing small business accounting service provider, and Mazuma’s co-founder and CEO, Lucy Cohen (a qualified accountant), recently published a paper, “Probably Wrong”, that raises concerns about precisely this issue with the UK’s tax authority, HMRC, and its Making Tax Digital programme. MTD for Income Tax came into effect in April and mandates the use of technology in the now quarterly tax filing workflows of millions of the UK’s smallest microbusinesses, which is all well and good. However, some of the commercial apps hurriedly built for this purpose claim to utilise LLMs to automate and propagate transaction categorisation, some without human-in-the-loop interfaces for an accountant or bookkeeper to cut in and check or safeguard the accuracy of the bookkeeping.

Lucy’s central argument is straightforward: accounting data is not like other data. Every figure in a set of accounts must be traceable to its source and independently verifiable, and a record that cannot be shown to be correct is legally deficient regardless of whether it happens to be numerically accurate. An AI system that generates figures through pattern matching rather than by applying accounting rules to evidence cannot satisfy that standard, and the error may be invisible until it surfaces in front of an auditor or HMRC, or indeed it may remain invisible permanently, yet the error and its associated monetary deficit sit on the ledgers of either the business or HMRC.

I think the deepfake parallel also holds here. AI-generated accounts can be convincing until the moment they are relied upon for something that matters, at which point the absence of a sound foundation becomes visible — and the cost of an accounting error is not bounded by the size of the mistake at the point it was made, but by the reach of the financial system into which it propagated.

Lucy makes a number of points and helpful recommendations in her paper that I hope are taken seriously, not least asking HMRC to require developers to tag transactions that have been exclusively AI-handled and categorised, which reminded me of this old tweet.

At least in the sphere of accounting, thanks to Lucy, the AI accounting problem now has a name, Stochastic Contamination, and she poses a legitimate and fair challenge to legislators, and I hope one that prompts genuine reflection and debate.

A natural consequence of large numbers of new products flooding into a category is noise, and in a noisy market, established distribution and brand trust become more valuable, not less. The incumbents who already have brand equity, customer relationships, trust, and embedded workflows are sitting on assets that have quietly appreciated while everyone else has been focused on what AI can build. Channel and installed base were always moats. They are deeper moats now than they were before the noise began. So, that’s at least one point to established vendors like Xero, Sage & Intuit still grappling with the SaaS-pocalypse narrative.

But noise is only part of the problem facing the current wave of AI-assisted challengers. There is a deeper question that buyers of business software have always asked, usually without articulating it: Is this company real?

Not real in the legal sense, but real in the sense of having organisational substance: a team with depth, processes that continue to function when the founder is on holiday, a support function that actually answers, a product roadmap that reflects genuine understanding of the problem rather than the enthusiasm of someone who once had the problem themselves, and a balance sheet that gives reasonable confidence the lights will still be on next year. These are not glamorous qualities, and nobody puts them in a pitch deck, but they constitute the invisible infrastructure of trust that matters enormously when a business is considering handing over its operations, its invoicing, its payroll, its customer data, its financial records, to a software product it will depend on daily.

AI-assisted coding has lowered the barrier to building the product without lowering the barrier to building the company, and for most buyers in most business software categories, the company is at least as important as the product, and often more so.

This is far from conjecture. Merete Hverven, CEO of Visma, one of Europe’s largest business software companies serving 2.5 million SMB customers, recently published findings on LinkedIn, “In the AI era, trust is the ultimate software feature”, from a study of nearly 2,000 small business leaders on what drives their willingness to choose and pay for business software. The results cut against most of the current market narrative in ways that are worth taking seriously. “Safe and secure” and “reliable” both featured in the top five purchasing criteria. “AI-powered” landed second to last.

Hverven’s interpretation goes considerably further than the headline number alone suggests. “Our customers’ ranking of AI,” she wrote, “is a sophisticated risk assessment,” meaning they are not rejecting AI from ignorance or technophobia but expressing a considered and entirely rational view that the reliability of the tool comes before the intelligence of the tool, and that the trustworthiness of the company behind it comes before either, particularly in the mission-critical world of accounting and payroll where the cost of an error is not a minor glitch but a legal liability or a broken relationship with the employees depending on you. The florist in Berlin and the plumber in Warsaw are not waiting to become prompt engineers. They are waiting for the tools they already depend on to become reliably, invisibly better, and that is a meaningfully different product requirement from the one most of the current generation of AI-assisted challengers are building toward.

This possibly explains why I winced a tiny bit when established SaaS category leaders, Xero (disclosure: I hold Xero stock) and Intuit, were falling over themselves to announce strategic partnerships with Anthropic the other week.

Not because I think it’s a bad idea to partner with Anthropic, and they have to start somewhere, but partly because incumbents sometimes fall into the trap of treating huge platform shifts as features.

In the summer of 1995, I was a fresh-faced young account manager working at the PC-era British desktop accounting software vendor Pegasus, when we shipped an update to customers that added a single field to the customer and supplier contact records to store the contact’s email address. This tiny change meant that invoices and statements could be emailed to customers as PDFs instead of being sent by post or fax. The communication that accompanied this update went something along the lines of, “We’ve internet-enabled accounting!” — and while this was technically 1% correct, now with the benefit of three decades of hindsight, it was also remarkably, if innocently, short-sighted.

But I winced mostly because communicating the Anthropic partnership broadly to everyone drives messaging that would otherwise be intended for a technical, specialist developer audience (including developers inside the customer businesses) through a general mainstream channel that’s neither attuned to nor likely to want to hear it.

You can see why they did it — they’re speaking to investor analysts through their customers to attempt to rebut the SaaS-pocalypse narrative, and that is understandable.

But it’s bringing the equivalent of a trade-counter conversation to the retail shop floor, and introducing something the vast majority of customers aren’t qualified to use or understand, risking confusion about the value exchange, and, worse, it could erode trust if it backfires.

The irony of this moment is that the technology works well enough to produce convincing, if wafer-thin ‘products’, sets of accounts, and investor narratives, but not yet reliably enough for any of them to be taken at face value.

That is, in essence, what a deepfake era looks like from the inside.

Sources and references

Read the original on everscale.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.