RSS Amplifier

AI regulation, standards and reality · Jun 30, 2026

prEN ISO/IEC DIS 23282: How to measure whether Natural Language AI is accurate

0
Sign in to vote or save

This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.

What the new NLP evaluation standard covers, and why providers of high-risk and general-purpose AI should read it

This article covers prEN ISO/IEC CD 23282 — Artificial Intelligence: Evaluation methods for accurate natural language processing systems. It is being developed by ISO/IEC JTC 1/SC 42 in parallel with CEN-CENELEC JTC 21, led by Lauriane Aufrant. It is at Draft International Standard / Enquiry stage and potentially subject to substantial change. Do not cite it as a current standard.

Natural language processing is where the current iteration of AI has taken us. Large language models, chatbots, document triage, translation, summarisation, voice assistants, code generation. The systems that now sit inside recruitment, credit, medical, and public-service workflows and many read or write natural language at some point. Most of the AI systems that worry regulators today are, at some level, language systems.

Yet until now there has been no harmonised methodology for how you actually measure whether a natural language system is accurate. Providers pick their own metrics, apply them inconsistently, and report numbers that cannot be compared across implementations. A year ago, in my article on task-level AI assurance, I described ISO/IEC 23282 as a “much-needed and in-progress” standard. There is now a draft to read (well, soon).

It is worth your time if you work with NLP.

Read more

Read on adamleonsmith.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.