AI
How to Evaluate an AI Vendor When You Are Not an Engineer
Big Sky Consulting Group · August 21, 2026 · 7 min read

Every demo works. That is the problem
You have sat through four AI pitches this quarter. All four demos ran clean. All four vendors used the same three words. And you walked out of each one without a way to tell whether you had just seen a product or a prompt.
That is not a failure of your technical judgment. It is the intended outcome of a well-built demo. Demos are run on data the vendor chose, in a scenario the vendor rehearsed, with the vendor's best engineer sitting three feet from the keyboard. Nothing in that room is a test. The test is what happens on your data, in your process, on a Tuesday, when the person operating it is the coordinator who has been doing this job by hand for six years and is not thrilled about the change.
We sit on the buying side of that table. We do not build this software, which is the entire reason our read on it is worth anything. What follows is not a procurement checklist you can run yourself, because the questions that decide the outcome depend on answers only your systems can give. It is the shape of the problem, and the small number of questions that reliably separate the vendors worth diligence from the ones that are not.
The market is currently rewarding the wrong thing
Two numbers are worth carrying into every one of these meetings.
Gartner reported in June 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, on costs, unclear value, or missing risk controls. In the same research they named the mechanism driving a good share of it: agent washing, the rebranding of existing assistants, RPA, and chatbots as agentic without the underlying capability. Their estimate was that of the thousands of vendors positioning as agentic AI, roughly 130 were real.
The second number comes from MIT's State of AI in Business 2025, which reviewed 300 public deployments alongside interviews and surveys of executives. Ninety-five percent of the pilots studied produced no measurable P&L impact. Buried in that report is a finding that cuts against the instinct most operators have when they get burned: externally built tools succeeded roughly twice as often as internal builds. The problem is not that you bought. It is what you bought, and what you did not ask.
Read those together and the picture is not that AI does not work. It is that the category has more sellers than substance right now, and the filtering job has been quietly moved onto the buyer. Vendor marketing will not do it for you, and neither will a peer's reference call, because the peer is eleven months into a deployment they are not yet willing to call a mistake.
The four questions
We open every vendor review with variations of the same four. They are not technical questions. Not one of them requires you to know what a context window is.
What happens when it is wrong?
Every one of these systems is wrong sometimes. A vendor who cannot describe their failure rate in your specific use case, in a number, has not measured it. More telling is the second half: what does the system do when it is uncertain? The good answers are boring and operational. It flags for review, it routes to a human queue, it refuses below a confidence threshold and logs why. The bad answer is a smooth reassurance about accuracy improving over time. Ask what the review queue looks like on day one and who staffs it. That question has ended more evaluations for us than any security review.
What is this on top of?
Most AI products in this market are an interface, a set of prompts, and some integration work sitting on a model licensed from someone else. That is not disqualifying. Some of the best tools we recommend are exactly that, and the integration is where the real work is. What matters is whether the vendor will say so plainly. When a founder answers this question directly, describes which model they use, why, and what happens when that provider deprecates it, you are talking to someone who has thought about running a business rather than raising one. Evasion here correlates almost perfectly with the thin-wrapper case, where your renewal price is set by a provider the vendor cannot negotiate with.
Who owns the data, in writing?
Not the answer in the meeting. The clause in the agreement. Confirm that your inputs and the outputs both remain yours, and that neither is licensed back for training the vendor's models. Confirm what happens to it if the company is acquired or shuts down, which in this category is a live question rather than a theoretical one. And be aware that a SOC 2 report, which most buyers now treat as the bar, tells you about the vendor's security controls. It tells you almost nothing about model behavior, training data provenance, or what the system does with a record it was not supposed to see.
What breaks in the process before the software matters?
This is the one we care about most, and it is the reason a good portion of our vendor reviews end with a recommendation not to buy anything. If the underlying process has three undocumented exception paths and two people who handle them by memory, no vendor on your shortlist will fix that. They will automate the documented path, the exceptions will keep arriving, and in month five you will have both a subscription and the same two people. We wrote about the same pattern in the pilot context in why AI pilots stall short of business outcomes, and it repeats at the vendor selection stage almost word for word.
The regulatory floor moved this month
One change worth knowing about, because a vendor who is unaware of it is telling you something.
The EU AI Act's Article 50 transparency obligations became applicable and enforceable on 2 August 2026. They cover systems that interact with users, generate text, images, audio, or video, or infer emotion and biometrics, and they apply whether or not the system is classed as high risk. Disclosure that a user is talking to an AI, marking of AI-generated content, deepfake labeling. Systems already on the market before that date have until 2 December 2026 to meet the machine-readable marking requirement. Separately, the AI Omnibus that came into force on 27 July 2026 pushed most high-risk obligations back to 2 December 2027, though anyone deploying AI in recruitment, evaluation, or worker monitoring should keep preparing on the original timeline.
If you have European customers, staff, or entities, this is your vendor's problem and it becomes yours on signature. Ask how they handle Article 50. A vendor with a clear answer has a compliance function. A vendor who has to check has told you the size of their company more honestly than their website did.
What we do not put in an article
Where the shortlist gets cut is in the specifics: which of your processes has enough volume and enough stability to justify a tool at all, what the switching cost looks like in eighteen months, how the pricing model behaves when your usage triples, and which two of the four vendors are actually competing for the same job. Those depend on your systems and your numbers. We can tell you which questions decide the outcome. We cannot tell you your answers from here, and any firm that offers to is selling the same certainty the vendors are.
The related work matters too. If you are weighing spend across several initiatives at once, the framing in measuring return on AI investment applies well outside education. If the choice in front of you is architectural rather than procedural, building a reusable AI stack covers why the composable option usually survives the second vendor change.
One more thing before the meeting. The best signal in the room is a vendor who tells you their product is wrong for one of your use cases. It is rare, it is free to give, and almost nobody gives it. When you hear it, write it down. That vendor just did the buy-side's work for you, which is either integrity or an unusually good closing technique. Either way it is data, and around here we consider that a fair trade.
If you have three quotes on the desk and no defensible way to rank them, that is the conversation to have before the quarter closes rather than after the contract is signed. Bring us the shortlist. We will tell you which one to buy, or that the answer is none of them and the money belongs in the process instead.
