I’ve spent fifteen years on the buyer’s side of vendor evaluations, and the closing ritual is the same everywhere. The sales engineer sends over the SOC 2 report. Someone in security opens the PDF, checks that the auditor’s opinion is clean, and marks the review complete in the procurement system. The deal moves.
The ritual deserves respect. It worked for a generation because of an assumption so safe nobody bothered to write it down. Software behaves the same at your company as it does everywhere else, so one audit could stand in for everyone’s testing.
AI vendors just broke that assumption, and the break is invisible from inside procurement, because the paperwork still shows up on schedule. It’s the wrong paperwork. MIT Sloan and BCG surveyed more than 1,240 executives across 59 industries and found that 78 percent of organizations use third-party AI tools, and 55 percent of all AI failures come from them. More than half the failures are coming from the place the binders say is covered.
What the audit actually says
Let me be fair to the paperwork before I take it apart, because a SOC 2 is a real thing with real value. An accountant examines a vendor’s controls against criteria for security, availability, processing integrity, confidentiality, or privacy and issues an opinion. Microsoft’s own compliance page says the opinion covers whether the vendor’s controls were designed appropriately and were “operating effectively over a specified time period,” a period that gets examined after it ends, with the report usually issued months later.
Read that as a buyer and two words should stick: theirs, and period. The report describes the vendor’s house. Their access controls, the locks on their doors. And it describes a stretch of time that was already over before your contract started. That’s useful for judging whether the vendor runs a serious shop. In my experience it’s also where the scrutiny ends. I can’t remember a single deal that stopped over something inside one. The report gets skimmed for the opinion letter, filed, and never opened again, and it says nothing about what their model will do inside your business next quarter. That was the question all along.
You’re reading last year’s weather report for somebody else’s zip code.
The first product that behaves differently at your house
The old assurance model made sense because software was deterministic. Same input, same output, at your company and at every other one, so a test anywhere was a test everywhere. A model is a different animal. What it produces depends on your data, your prompts, your customers, and the particular strange questions your business attracts. Nobody can certify in advance how their model behaves in contact with your operation, because that behavior doesn’t exist yet. It gets created the day you connect the two.
Air Canada found out where that leaves a buyer. Its website chatbot invented a bereavement refund policy for a grieving customer, contradicting the real policy on a page the bot itself linked to. The airline argued before a tribunal that the chatbot was “a separate legal entity that is responsible for its own actions,” lost, and paid $650.88 in damages. I wrote about the case in the software workforce post, and the detail that matters here is the timing. Whatever review Air Canada ran before launch, the chatbot passed it. The failure wasn’t available for inspection until a real customer with a real funeral asked the question.
The takeaway: AI risk lives in the pairing of their model with your business, and a pairing can’t be pre-certified. The proof you need can only be produced at your house, after you sign.
The product keeps moving after you sign
Researchers at Stanford and UC Berkeley ran the same questions through GPT-4 in March 2023 and again in June, and published the differences. Accuracy at answering whether numbers were prime fell from 84 percent to 51 percent. The June version was also less willing to answer sensitive questions and made more formatting mistakes in code. Princeton’s Arvind Narayanan and Sayash Kapoor pushed back on the headline, arguing the model’s underlying capability hadn’t declined, only its behavior. They also wrote the sentence every buyer should frame. The fine-tuning these models regularly undergo “can have unintended effects, including drastic behavior changes on some tasks,” and when behavior shifts, workflows built on the old behavior “might stop working.” The fight was about why. Nobody in it disputed that the thing changed underneath the people using it.
From the vendor’s side, this is routine maintenance. In April 2025, OpenAI shipped a personality update to GPT-4o, watched it turn “overly flattering or agreeable,” their words, and rolled it back within days. The postmortem says they “focused too much on short-term feedback.” Credit where due, they published the whole thing. Now check your contract for the clause requiring your vendor to tell you when the model behind your customer service gets updated, goes sideways, and gets rolled back. I’ve read a fair number of AI agreements by now, and not one of them has had that clause in the vendor’s standard paper.
Retirement runs on a schedule too. OpenAI’s deprecation policy promises at least six months’ notice before shutting down a generally available model, three months for specialized variants, and reserves the right to retire preview models on about two weeks’ notice. The version your team spent a quarter evaluating carries a shutdown date, and the vendor picks it.
The takeaway: the product you’re running is no longer the product you evaluated. An AI vendor evaluation expires like milk, so put re-testing on the calendar right next to the renewal date.
The burden of proof crossed the table
Step back and look at what the whole assurance apparatus is for. Audits and questionnaires exist so that ten thousand buyers don’t each have to test the same product. The vendor proves the product once, an accountant vouches, and everyone else reads the report. The efficiency rests on two assumptions, and AI just broke both. A model doesn’t behave the same everywhere, and it doesn’t stay what it was.
Which puts buyers somewhere new. For the first time since companies started buying software, the meaningful testing can only happen on your side of the table, on your own cases, for as long as the contract runs. The vendor can’t produce the evidence you need, no matter how cooperative they are, because the evidence only exists where their model meets your work.
The compliance industry has noticed the gap and is answering with the tool it knows how to make: a new certificate. ISO 42001 is the one your vendors will start bringing up, if they haven’t already. AWS became the first major cloud provider certified under it in November 2024, and the standard covers, in its own framing, “requirements and controls for organizations to promote the responsible development and use of AI systems.” That certifies management processes. It’s their house again, better organized this time. Ask for it anyway, since a vendor with adult supervision beats a vendor without it. Just don’t file it as a promise about how the model behaves on your work. No certificate anywhere can make that promise, and that isn’t ISO’s fault. The thing being promised doesn’t hold still long enough to certify.
The takeaway: for a generation, the vendor proved the product so you didn’t have to. That arrangement is over for AI. Testing is now a buyer-side job that never finishes, and at most companies it’s a job nobody holds.
Nobody owns the watching
Ask a leadership team who owns AI vendor behavior after go-live and watch the answer travel. Procurement points at security, security points at the business unit, and the business unit points at the vendor, because that’s what the subscription is for. I’ve sat in the conference rooms where that circle forms, and nobody in them is being lazy. Each person is right that it isn’t their job. Everyone owns some part of the purchase, and the watching belongs to no one.
This rhymes with something I wrote earlier this year, you employ more software than people, and most of that software has no manager either. An AI vendor is a contractor working inside your business every day. Accepting a SOC 2 and filing it is hiring the contractor because the background check cleared, then never once reviewing the work. A background check is not a performance review.
The fix costs less than the tools do, and it fits on two lists. First, the contract.
Terms That Buy Visibility
- Change notice. Written notice before the model version serving your account materially changes, far enough ahead to re-test. Vendors already announce retirements to the public on their own schedule. Your agreement can require they notify you, directly, with dates.
- A pin with a runway. The right to stay on the version you evaluated for a defined window, and a dated migration period when it retires. Big labs already publish date-stamped model versions, so the clause simply makes that stability yours.
- Eval rights. Written permission to benchmark the service on your own cases and to share the results with your executives and your board. If anything in the standard terms restricts benchmarking, strike it at renewal and watch how hard they fight.
- Rollback disclosure. When the vendor rolls back or hotfixes the model behind your product, you get notice in writing within days, whether or not it made the news. OpenAI told the entire internet about its rollback. Your vendor should at minimum have to tell you.
- Exit with your records. Prompts, evaluation sets, fine-tuning data, and logs, in a format you can hand to the next vendor. The denominator post covered why exit rights protect your price. They also protect your memory. A documented history of what the model did is how the next vendor gets up to speed, and how your lawyers reconstruct events if it ever comes to that.
Second, the job. Give one named person the watching, and put it in their goals. A committee is how the circle re-forms. What the owner actually does would look familiar to anyone who has ever managed a contractor:
The Owner’s Quarterly Loop
- Baseline before go-live. A fixed set of real cases from your business, run through the tool, scored, dated, saved. Fifty cases is plenty to start.
- Re-run on a calendar and on every change notice. The score is your drift alarm. If accuracy on your cases slides the way it did in the Stanford study, you find out from your own dashboard instead of from a customer.
- Log the differences where vendor risk can see them. A one-page quarterly readout of what changed, what it broke, and what it cost.
- Rehearse the bad day before it comes. Know today whether you can pin the version, roll back, or fail over to the second vendor you kept warm. A fallback nobody has rehearsed is a wish.
- Re-price annually. Behavior drift, retirement schedules, and renewal quotes belong in one conversation. Walk in with your eval history and you’re negotiating with data instead of with hope.
None of this is exotic. It’s how companies already manage contractors, parts suppliers, and everyone else whose output changes over time. AI vendors got a temporary exemption from that discipline because their product arrived wearing software’s old clothes, and software had spent thirty years earning the right to be trusted on paperwork.
I still read every SOC 2 that crosses my desk. It tells me whether the vendor runs a serious shop, and I want to know that. I ask for the new certificates too. Then comes the part no binder has ever done for anybody. Somebody watches the work itself, on real cases, with their name on the job.
The paperwork tells you about their house. The risk lives in yours.
How RLK Can Help
The Vendor and Spend Strategy engagement writes the terms above into your actual contracts, RFPs, and renewals, next to the pricing and exit protections that belong beside them. The Board Readiness engagement builds the technology story your board actually needs, including who owns which risk. And my free AI Diagnostic shows you in about eight minutes where your AI program stands, including the parts nobody is watching. Start the conversation.
Sources
- MIT Sloan, “Third-party AI tools pose increasing risks for organizations”
- AICPA, SOC 2 examinations of controls at a service organization
- Microsoft Learn, SOC 2 Type 2 compliance offering
- Chen, Zaharia, and Zou, “How Is ChatGPT’s Behavior Changing Over Time?” (arXiv)
- Narayanan and Kapoor, “Is GPT-4 getting worse over time?”
- OpenAI, “Sycophancy in GPT-4o: What happened and what we’re doing about it”
- OpenAI, API deprecations and model retirement notice periods
- AWS, ISO/IEC 42001 accredited certification announcement
- American Bar Association, “BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot”
- CBC News, “Air Canada found liable for chatbot’s bad advice”