GovernanceCore
Tools & PlatformsIntermediate

Third-Party and AI Vendor Risk Assessment

Buying an AI system does not buy out the obligation. This is how to assess an AI vendor: the questions that produce evidence rather than assurances, the contract terms that matter, and why the assessment cannot be a one-off.

AI Governance TeamPublished August 3, 202614 min read
Key takeaways
  • Deployer obligations do not transfer with a purchase. Under the EU AI Act you carry your own duties regardless of who built the system.
  • Article 25 allocates responsibilities along the AI value chain, which is the hook for contractual flow-down to your suppliers.
  • The distinguishing risk of AI vendors is that the product changes after you buy it, often without a deployment on your side.
  • Ask for evidence you can inspect rather than assertions you have to trust, and score the answer on what was produced.
  • OMB M-25-22 is a published benchmark for what a very large, risk-averse buyer now requires, and private buyers can borrow the structure for free.
Art. 25EU AI Act clause allocating responsibilities along the AI value chain, the basis for contractual flow-down
Art. 26Deployer duties that survive outsourcing. Buying a system does not buy out the obligation
30 Sep 2025Date from which OMB M-25-22 requirements attach to US federal AI solicitations
0Number of vendor assurances that constitute evidence on their own

Third-party risk management is a mature discipline. Most organisations have a questionnaire, a tiering model, and a review cadence. Then AI arrives and the questionnaire asks whether the vendor holds ISO 27001 and encrypts data at rest, which tells you something about their security posture and nothing about whether their model behaves.

The gap is not that AI vendors are riskier. It is that three assumptions underneath classic vendor assessment stop holding.

The product changes after you buy it. Software you procured last year behaves the same today unless someone deployed. A hosted model can change underneath you without any deployment on your side, and the vendor may not tell you.

You cannot inspect the thing you are buying. You can review an architecture diagram. You cannot read a model's weights and conclude anything useful about its fairness on your population.

The obligation does not transfer. This is the one that surprises people. Under the EU AI Act, a deployer carries duties of its own regardless of who built the system. Outsourcing the build does not outsource the accountability.

This article is about assessing any AI vendor. If your question is which governance platform to buy, that is a different exercise and we cover it in best AI governance platforms.

Scoping: which vendors are AI vendors

Start here, because the answer is larger than the list procurement currently holds. Three categories are routinely missed.

  • Embedded features. The CRM that added a summarisation feature, the ticketing tool with suggested replies, the HR platform that now scores candidates. Nobody bought AI. AI arrived in a release note.
  • Sub-processors of models. Your vendor built the application and calls someone else's model. Your exposure includes a party you never contracted with.
  • Tools bought below the procurement threshold. The team expensing a licence on a card. This is the same discovery problem as shadow AI, and it needs the same techniques.

A practical trigger to add to your intake: does this product process our data or our customers' data with a model, or produce an output a person or a system will act on. If yes, it is in scope regardless of how it was bought or what the contract calls it. Feeding the result into the same register as your internal systems is the point of an AI model inventory.

The questionnaire, organised around evidence

The failure mode of vendor questionnaires is that they collect assertions. A vendor writes yes, you file it, and neither of you has learned anything. Reframe every question so the answer is a document, a result, or a demonstration.

AreaAskWhat good evidence looks like
Model provenanceWhich models does the product use, including sub-processors, and which versionsA named list with versions, and a statement of which are third-party. Vagueness here predicts vagueness everywhere
Training dataWhat categories of data was it trained or fine-tuned on, and on what legal basisDocumented categories and provenance. For GPAI, the summary the AI Act contemplates
Our dataIs our data used for training, by default or on opt-in, and for how long is it retainedThe contractual clause, not the marketing page. Check the default
EvaluationHow was performance measured, on what population, and what were the resultsAn actual evaluation report with metrics and a described test set. Absence is itself a finding
FairnessHas it been tested for disparate outcomes, on which attributes, with what resultA disparity table. If they cannot produce one, you will have to run your own. See bias auditing and fairness testing
Adversarial testingHas it been red-teamed, by whom, against what threat modelA report with findings, severities, and fixes. See AI red-teaming
Change controlHow will you tell us the model changed, and how far in advanceA written notification commitment with a period. Most vendors have none until asked
Human oversightWhat is the product designed to assume a human will checkA clear statement of intended use and known limitations, which the AI Act expects providers to supply
Incident dutiesHow fast will you tell us about an incident, and will you cooperate with our reportingA period short enough to let a two day regulatory path be met. See AI incident response
SecurityStandard controls, plus model-specific: prompt injection, tool permissions, tenancy isolationCertifications for the general case, test evidence for the model-specific case
Compliance postureWhich role do you take under the EU AI Act, and for GPAI have you signed the Code of PracticeAn explicit role claim. A vendor unsure whether it is a provider has not done the work
ExitWhat happens to our data, and can we get our configuration and history outAn export path you have seen work, not a clause promising one

The role question deserves emphasis. The EU AI Act assigns duties by role, and a vendor that cannot say whether it considers itself a provider or is acting as your deployer has not thought about the regime. That answer also tells you which duties are theirs and which remain yours.

On general-purpose models, the GPAI Code of Practice is a reasonable thing to ask a foundation model provider about. Signing is voluntary, but it is one route to evidencing AI Act compliance for GPAI models, and a non-signatory should be able to explain how it evidences compliance instead.

Evidence over assertion

A short rule that changes outcomes: score the answer on what was produced, not on what was claimed.

ScoreMeaning
3Produced a document or demonstration we could inspect, and it addressed the question asked
2Produced something partial, or a document that addressed an adjacent question
1Asserted a practice with no artefact
0Could not answer, or answered by pointing at a certification that does not cover the question

Two patterns to watch. A security certification offered as an answer to a model behaviour question is a category error, and a common one. And a vendor that answers every question fluently but produces nothing is often a reseller who does not control the model. Neither is disqualifying by itself, but both change what you have to do yourself.

Contract terms that matter

Assessment findings only become controls if the contract carries them. Article 25 of the EU AI Act deals with responsibilities along the value chain, which is the mechanism for allocating duties between you and your suppliers. In practice, six terms do most of the work.

1. Role and duty allocation
State each party's role under the applicable regime and which obligations sit where. Ambiguity here is what produces a dispute during an incident.
2. Notice of material model change
Define material, and set a notice period. Without this, your system's behaviour can change with no deployment and no warning, and your change control never fires.
3. Incident cooperation with a clock
A notification period, an obligation to provide the technical facts you need for your own reporting, and a named contact. The period has to be short enough for the tightest regulatory path that could apply to you.
4. Evidence and audit rights
A right to receive evaluation and testing evidence on a cadence, not only at onboarding. Full audit rights are hard to win from large vendors; a right to documentation and to results is usually achievable and often sufficient.
5. Data use and training
Whether your data trains their models, with the default stated explicitly rather than inferred, plus retention and deletion on exit.
6. Sub-processor transparency
A current list of model providers behind the product, and notice before it changes. Your exposure runs through parties you did not contract with.

Note one limit worth knowing about if you operate in Colorado. Under the state's automated decision-making statute, a contractual provision that indemnifies or holds a party harmless for its own acts or omissions in a Colorado anti-discrimination claim arising from a covered system is void as contrary to public policy. Indemnities are not a route around your own conduct there.

Tiering, so the depth matches the exposure

Running the full assessment on every vendor guarantees it gets run properly on none. Tier on consequence rather than on contract value.

TierTriggerAssessment depthReassessment
CriticalOutput materially influences a consequential decision about a person: employment, credit, housing, insurance, healthcare, education, essential servicesFull questionnaire, evidence inspected, independent testing on your population where feasible, governance forum approvalAnnual, plus on any material model change
HighCustomer-facing at scale, or an agent that can take real actions, or processes special category dataFull questionnaire with evidence for evaluation, security, and change controlAnnual
ModerateInternal productivity with a human checking every outputShort-form: data use, retention, security, model provenanceAt renewal
LowNo sensitive data, no consequential outputRegister it. Confirm data handlingAt renewal

The critical tier deliberately mirrors the domains that consequential-decision statutes tend to enumerate, because those are the systems where a vendor's failure becomes your regulatory problem.

Continuous, not a one-off

This is the part most programs skip, and it is where AI vendor risk genuinely differs. Four triggers should reopen an assessment:

  • The vendor changes the model. Which you will only know if you contracted for notice.
  • Your use expands. A tool approved for drafting is now feeding a decision. The tier changed even though the vendor did not.
  • Renewal. The natural gate, and the moment you have the most bargaining power for adding terms you did not win first time.
  • An incident anywhere. Including at another customer. A public failure involving a model you use is a reason to look at your own logs.

Monitoring the vendor's own behaviour matters too. Drift in a hosted model is invisible from the outside unless you are measuring outputs on a stable test set of your own. A small held-out set, scored on a schedule, is the cheapest early warning available for bought AI.

The public sector benchmark

If you want to know what a very large, risk-averse buyer requires, read what one publishes. OMB M-25-22, issued on 3 April 2025 and replacing the earlier acquisition memorandum, governs how US federal agencies buy AI. Its requirements attach to solicitations issued on or after 30 September 2025 and to renewals or extensions exercised on or after 1 October 2025, so it reaches existing vendor relationships at their next decision point.

For a private buyer it is a free, maintained baseline covering performance, risk management, competition, data rights and lock-in. Borrowing its structure also has a practical benefit: vendors selling into government are already preparing to answer it, so your questions land on documents that exist. The companion memorandum on federal AI use is covered in AI governance in the public sector.

Red flags

  • Cannot name the models behind the product, or names them only in general terms.
  • No evaluation evidence, and no plan to produce any.
  • Refuses to commit to notice of material model change.
  • Offers a security certification as the answer to a model behaviour question.
  • Claims regulatory compliance without being able to state which role it occupies.
  • Contract silent on whether your data trains their model, with a permissive default buried in the terms of service.
  • No export path, or one nobody has tested.
  • Answers improve markedly when you ask for the artefact, which suggests the first answers were aspirational.

Frequently Asked Questions

Does buying an AI system transfer the compliance obligation to the vendor?

No. Under the EU AI Act, duties are allocated by role, and a deployer has obligations of its own regardless of who built the system, including on human oversight, use in accordance with instructions, and informing the provider of a serious incident. Article 25 lets you allocate responsibilities along the value chain contractually, which shapes who does what between you and your supplier, but it does not remove your role's duties. The practical consequence is that procurement becomes a governance control rather than a purchasing step.

What if the vendor refuses to provide evaluation evidence?

Treat it as a finding with three possible responses rather than an automatic disqualification. Test it yourself on your own population, which is the strongest option and sometimes the only credible one for a consequential decision. Constrain the use case so the missing evidence matters less, for example by requiring human review of every output. Or decline. What you should not do is record the refusal and proceed as though the evidence existed, because that is the version an auditor will find.

How do we assess a vendor that just calls a large model API?

Assess both layers, and be explicit about which risks sit where. The application vendor owns the prompt design, retrieval sources, guardrails, tool permissions, logging and their own change control, and all of that is assessable normally. The model provider owns training data, base model behaviour and versioning, and you reach it indirectly. Ask for the sub-processor list, ask what happens when the underlying model version changes, and confirm whether your data reaches the model provider and on what terms.

Our procurement questionnaire already covers security. What do we add?

Six additions cover most of the gap: model provenance including sub-processors, training data and whether your data is used, evaluation and fairness evidence with results rather than a claim, change control with a notice commitment, incident cooperation with a period, and a clear statement of intended use and known limitations. Keep the existing security section, and extend it with the model-specific items such as prompt injection resistance and tool permission scoping.

How do we find AI that entered through features we already bought?

Three methods, used together. Read release notes for your top vendors, since this is how most embedded AI arrives. Query expense and SaaS management data for new tools and for AI-adjacent line items. And ask teams directly what they are using, which works better than it should because most people are not hiding anything, they simply did not consider a feature to be a procurement event. Then register what you find and tier it.

How often should critical AI vendors be reassessed?

Annually as a baseline, and on trigger rather than only on calendar. The triggers that matter are a material model change, an expansion of how you use the system, renewal, and any incident involving that vendor or model even at another customer. For anything in the critical tier, add your own periodic measurement on a held-out test set, because a hosted model can drift without any notification and without any change on your side.

ai-vendor-riskthird-party-riskdue-diligenceai-procurementsupply-chainEU-AI-ActGPAIcontract-termsvendor-tieringmodel-change
AI Governance Team
Editorial Team

Expert analysis and in-depth reporting from the AI Governance Core editorial team, covering enterprise AI compliance, ethics, and responsible AI practices.

Related analysis

Choosing an AI Governance Platform: What to Evaluate
Tools & PlatformsIntermediate

Choosing an AI Governance Platform: What to Evaluate

AI Governance Team··13 min read
How to Build an AI Model Inventory (and Why It's Step One)

How to Build an AI Model Inventory (and Why It's Step One)

AI Governance Team··11 min read
Building an AI Governance Operating Model

Building an AI Governance Operating Model

AI Governance Team··13 min read