This article covers prEN ISO/IEC CD 23282 — Artificial Intelligence: Evaluation methods for accurate natural language processing systems. It is being developed by ISO/IEC JTC 1/SC 42 in parallel with CEN-CENELEC JTC 21, led by Lauriane Aufrant. It is at Draft International Standard / Enquiry stage and potentially subject to substantial change. Do not cite it as a current standard.
Natural language processing is where the current iteration of AI has taken us. Large language models, chatbots, document triage, translation, summarisation, voice assistants, code generation. The systems that now sit inside recruitment, credit, medical, and public-service workflows and many read or write natural language at some point. Most of the AI systems that worry regulators today are, at some level, language systems.
Yet until now there has been no harmonised methodology for how you actually measure whether a natural language system is accurate. Providers pick their own metrics, apply them inconsistently, and report numbers that cannot be compared across implementations. A year ago, in my article on task-level AI assurance, I described ISO/IEC 23282 as a “much-needed and in-progress” standard. There is now a draft to read (well, soon).
It is worth your time if you work with NLP.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.