It has been a while since I last got a chance to update you on progress at the intersection of Irish and AI; I hope to stick to a more predictable writing cadence over the next few weeks and months. So let’s start with the big news…
Back in early 2025 when I started working on this bench, there was only one LLM benchmark for Irish released, the UCCIX / IrishQA benchmark. It was a lack of knowing how good AI really was at Irish that inspired me to devise my own eval set to answer that question. And I wasn’t the only one thinking this: by the end of 2025, there were 4.
And now in 2026, thanks to Údarás na Gaeltachta, you can keep up with the latest releases and how they do as Gaeilge. See below for a snippet or here for the live leaderboard.
Ar scáth a chéile a mhaireann na daoine (strength comes from working together), and it is fantastic to be collaborating with researchers across the country to help understand the future of AI and the Irish language. On a separate note, I want to thank the readers of this Substack and my website for the support: as much as I do this to embody being caidéiseach (brimming with curiosity), your interest, questions and support are a big inspiration to continue working on this over the now years.
Now back to the regular programming—how are the models of 2026 doing?
Fable 5 lived up to the hype and has broken the glass ceiling of 80 that every model up until now has failed to break. All the frontier models are beginning to cluster around the 78-80% now as you can see from the model release vs accuracy graph and this zoomed-in bar chart of the top 10 models.
What this all means is that models at their baseline—no skills, tools, external memory or any harness—are becoming not just good, but very good. I think this is a watershed moment in approach and should encourage not just more LLM benchmarks, but new agentic benchmarks (see What’s Next?). From an end-user perspective, LLMs now likely outperform Gaeilgeoirí on grammar, if parallels can be drawn from the latest GaelEval for Scot’s Gaelic. That being said, the jagged edge of LLM and agentic knowledge and abilities persist.
So now that many of us will be using our model provider of choice for helping us work with Irish, one of the areas I wanted to look at was whether our results show models struggling with any particular type of linguistic facet of Irish. One of the most notable is that the Claude family of models struggle with lenition, with Fable, Opus 5 and Sonnet 5 all showing above average issues (up to 20% of their errors) knowing when to include or not include a h, for example, giving an-fuar instead of an-fhuar for ‘very cold’. This compares to other families showing 10-17% errors on lenition.
There will still be times you’ll want to open Gramadach Gan Stró or your favourite grammar book of choice. Particular care should be shown around the copula (which we also have problems with so its nice to know we are all in the same proverbial boat), dates and ordinal numbers. There is the most variance in model quality when it comes to the subjunctive (mean 50% accuracy, std. of 28%) and prepositional pronouns (std of 25%, though in this case the mean score is 80%).
The models base abilities, whether or not you allow thinking, don’t seem to change the results much, and can sometimes negatively affect the results as models “overthink” its pre-trained knowledge. I believe this is just a side-effect of the way we are benchmarking at the moment, and expect thinking, multi-step reasoning to greatly outscore the base models once they have actual tools to help them ground their (artificial) intellect.
Chinese models, generally open source models and our own Irish-language trained model, are performing better and better over time. We have moved so far so quickly on from the disappointments of the Llama 2 series and GPT OSS models. The one that jumps out is Qomhra, which sits in the centre of the pack, despite only being 8B parameters. There is so much promise into 2027 for open and local access to high-quality Irish Language LLMs.
Agents, agents, agents. The techno-optimism for the Irish language is growing. I am currently working on a smaller agentic Irish benchmark, with a continued focus on grammatical discrimination and generation. I hope to have some more updates soon ✌️

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.