RSS Amplifier

Competence Debt · Mar 17, 2026

The Vendor Evaluation Problem

0
Sign in to vote or save

Alex Blumentals · Competence Debt

I have been watching the same conversation play out in compliance teams across Europe for the past six months.

It goes like this: someone -- usually the DPO, sometimes the procurement lead -- realises that Article 26 of the EU AI Act requires them to verify their AI vendors’ compliance. They look at the timeline (August 2, 2026), panic briefly, then ask the obvious question: “Can’t we just upload the vendor’s documentation to ChatGPT and ask it to check?”

The answer is yes. And also no.

Yes, you can do exactly that. We have written the prompts for you. Today we published a blog post with six prompts -- one for each pillar of the Twin Ladder Standard -- that you can copy directly into any large language model. Upload your vendor’s transparency report, model card, DPA, and terms of service. Run the prompts. Score the results on a 0-3 scale. Calculate the percentage. Map it to the four Twin Ladder maturity levels.

We also provide three complete RFI templates -- 36 to 48 questions each -- free for registered users. Every question is mapped to a specific Article 26 deployer obligation. Every question includes scoring guidance.

This is not a teaser. There is no paywall. The prompts are in the post. The templates are on the platform. Use them.

Because the hard part was never asking the right questions.

I spent years inside procurement organisations. I have watched the information asymmetry between buyers and sellers play out across thousands of categories. The seller knows their product inside out. The buyer is spread across hundreds of purchasing decisions and cannot possibly match the seller’s depth in any one category.

AI vendor evaluation is this asymmetry at its most consequential. Your vendor has a compliance team producing polished transparency reports. You have a procurement analyst, a DPO, and a legal counsel -- none of whom have evaluated an AI system against the EU AI Act before.

The prompts help. They structure the inquiry. They ensure you ask about all six pillars. They catch vague language and missing documentation.

But they cannot do five things:

1. They cannot evaluate the vendor’s answers in the context of your organisation.

An LLM does not know that you are deploying this HR screening tool in Latvia, where the intersection of EU AI Act Article 26 and Latvian employment law creates obligations that differ from Germany. It does not know that your HR team has three people and no AI oversight function. The same vendor documentation produces a different compliance verdict depending on who is deploying, where, for what, and with what capabilities.

2. They cannot calibrate vendor requirements against your competence level.

A Level 1 organisation (Exploring on the Twin Ladder Scale) deploying a high-risk AI tool needs fundamentally different oversight controls than a Level 3 organisation (Implementing) deploying the same tool. The vendor’s documentation might be adequate for one and dangerously insufficient for the other. An LLM does not know your maturity level.

3. They cannot provide comparative benchmarks.

When we tell a client that a vendor scored in the 35th percentile for training documentation, that means something. It means we have enough evaluations to build a distribution. An LLM can tell you the documentation “appears comprehensive” -- it cannot tell you whether that is typical or exceptional.

4. They cannot produce legally defensible evidence.

If a regulator asks you to demonstrate due diligence under Article 26, a saved LLM conversation is not evidence. A Twin Ladder Vendor Compliance Report -- mapped to specific Article 26(1) obligations, scored against an open methodology, reviewed by a qualified assessor -- is designed to be evidence.

5. They cannot bridge from vendor evaluation to FRIA.

If your AI system is high-risk, Article 27 requires a Fundamental Rights Impact Assessment. Our platform takes the vendor evaluation data and pre-populates 30-40% of the FRIA. This is architectural -- it requires structured data flowing between the evaluation and the FRIA workflow. No prompt can replicate this.

We give away the methodology because we believe the framework should be a public good. The Twin Ladder Standard is CC BY-SA 4.0 -- open. If you can do this yourself, do it.

The paid service exists for the five things above: context, calibration, benchmarks, defensibility, and FRIA integration. These require a platform, accumulated data, and human expertise -- not a better prompt.

Read the full post with all six prompts, the scoring framework, and links to the free templates:

How to Evaluate Your AI Vendor in 30 Minutes (And Why It’s Not Enough)

The deadline is August 2, 2026. Five months. Your vendors have documentation. The question is not whether you can read it.

The question is whether you can evaluate it.

Alex Blumentals is the founder of Twin Ladder. The Twin Ladder Standard is an open framework (CC BY-SA 4.0) for assessing organisational AI competence across six pillars: Awareness, Policy & Data Protection, Training, Tools, Evidence, and Governance.

No posts

Read the original on competencedebt.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.